Tag
2 articles
Learn how NVIDIA's TensorRT Model Connect streamlines AI model deployment by converting Hugging Face checkpoints to optimized TensorRT engines in two commands, eliminating intermediate steps and enabling native C++ inference.
Learn to implement and evaluate a hybrid MoE-diffusion model that demonstrates the performance benefits of converting autoregressive LLMs into diffusion models for improved inference speed.